Øyvind Kolås [Fri, 8 Nov 2019 18:07:47 +0000 (19:07 +0100)]
build: opt out of unsafe math optimizations in reference and base
This makes the reference code paths used for verifying conversions in
extensions not involve for instance fast reciprocal approximations, see issue
#49. The extensions are still compiled with full optimizations.
Øyvind Kolås [Fri, 8 Nov 2019 12:54:15 +0000 (13:54 +0100)]
babl: adjust default BABL_TOLERANCE
This as a start of fixing issue #49, lower precision code gets generated
on AMD EPYC, due to use of rcpps to get a reciprocal - which has lower
precision on AMD EPYC. The conversions with AMD EPYC gets included with
a small margin - so we should not be shedding many other fast conversions
due to this.
Øyvind Kolås [Tue, 20 Aug 2019 00:54:51 +0000 (02:54 +0200)]
babl: register gray and gray alpha shortcuts between spaces
The formats "Y float", "YA float" and "YaA float" between different RGB
and gray spaces are all exactly the same. This registers memcpy shortcuts
for these combinations for all spaces. This avoids temporarily expand
into RGBA for many babl path fishes.
Ell [Mon, 19 Aug 2019 14:33:29 +0000 (17:33 +0300)]
avx2, two-table: don't segfault for NaN input
In the AVX2 and two-table linear-float => gamma-int8 conversions,
tweak the input bounds-check to handle NaN values. NaN would
previously lead to an out-of-bounds table lookup, and a segfault
(see issue #43).
Øyvind Kolås [Sun, 18 Aug 2019 23:00:11 +0000 (01:00 +0200)]
babl: reduce default max path length to 3 (5)
With the space invasion progressing, we're starting to have
many more conversions to search through - in some cases
looking for conversions - when there are no conversions
become too expensive with an exhaustive search of all
conversions up to 7 steps. (it is likely possible to reduce
the amount of conversions tried).
As well re-enable handling of grayscale ICC profiles - since
the formerly patological tests on gimp now pass.
Øyvind Kolås [Sun, 18 Aug 2019 22:01:13 +0000 (00:01 +0200)]
babl: add support for grayscale spaces
A grayscale space is just like an RGB space, but has a default set
of chromaticities for R,G,B. When ICC profiles are loaded only the TRC
is considered is considered (we could also include the whitepoint),
blackpoint tag is ignored (and should be baked into the used ICC
profile instead.)
This works with GIMP-2.10 but master of GIMP currently hangs when
trying to load a grayscale jpeg with attached grayscale ICC profile,
The if #0 on line 999 of babl/babl-icc.c needs to be turned into a
1 to enable grayscale icc profiles for further testing.
Øyvind Kolås [Fri, 16 Aug 2019 21:45:18 +0000 (23:45 +0200)]
babl: simplify logic in babl_epsilon_for_zero
We do not need to treate positive/negative avoided infinities differently,
the result is the same as long as we are consistent, and only using the
positive epsilon value leads to slightly simpler per-pixel conversion code
in both directions.
Øyvind Kolås [Tue, 6 Aug 2019 11:46:35 +0000 (13:46 +0200)]
meson: stop overriding libdir
Let meson do it's defaults, this mean we get installed in
prefix/lib/x86_64-linux-gnu/ instead of prefix/lib/
apparently which apparently is better - but different from
what autotools used to do; and leads to other paths also
needing potential adjustment.
Øyvind Kolås [Fri, 2 Aug 2019 11:53:44 +0000 (13:53 +0200)]
build: issue #41 builds without mmx/sse etc extensions do not work
The warnings about the SIMD flags specific variables not being set goes
away with this commit. Setting empty strings do not work since then gcc
ends up looking for input files that are the empty string, thus this patch
re-adds -Wall instead of setting an empty string.
This silences a false warning from gcc, where it doesn't recognize the type
of control flow occuring with the pre/post switches in the reference
conversion.
Add AVX2 conversions from Y float, YA float, RGB float, and RGBA
float, to Y' u8, Y'A u8, R'G'B' u8, and R'G'B'A u8, respectively.
The conversions use a lookup table, similarly to the two-table
conversions, indexed using AVX2's 256-bit gather instruction, which
allow us to process 8 floats at once. Over here, this conversion
is ~5x faster than the SSE conversions.
Note, however, that unlike two-table, we don't use a second
"feedback" table to correct the result, leading to an off-by-one
conversion error with a probability of ~0.1%, using a 2^16-element
table. This error rate is low enough for babl to use the
conversion, but might still be a bit too high regardless; it can be
further reduced with a bigger table.